HeteroRL (GEPO): Group Expectation Policy Optimization

HeteroRL is an architecture to decouple rollout sampling and policy optimization, the core of which is GEPO, an algorithm proposed to alleviate high variance of importance weights and training instability caused by increased KL divergence resulting from high latency due to heterogeneity in compute resources.

Date: 2026-09-05 Sat

Author: ArcaLunar